Low-Level Mark-Up and Large-scale LFG Grammar Processing
نویسندگان
چکیده
It is commonly believed that shallow mark-up techniques such as part-of-speech disambiguation or low-level phrase chunking provide useful information that can improve the performance of natural language processing systems, even those that ultimately require deeper levels of analysis. In this paper, we discuss three types of shallow mark-up: part of speech tagging, named entities, and labeled bracketing. We show how they were integrated into the ParGram LFG English grammar and report on the results of parsing the PARC700 sentences with each type of mark-up. We observed that named-entity mark-up improves both speed and accuracy and labeled brackets also can be beneficial, but that part-of-speech tags are not particularly useful.
منابع مشابه
Integrating a Unification-Based Semantics in a Large Scale Lexicalised Tree Adjoining Grammar for French
In contrast to LFG and HPSG, there is to date no large scale Tree Adjoining Grammar (TAG) equiped with a compositional semantics. In this paper, we report on the integration of a unification-based semantics into a Feature-Based Lexicalised TAG for French consisting of around 6 000 trees. We focus on verb semantics and show how factorisation can be used to support a compact and principled encodi...
متن کاملIntegrating Finite-state Technology with Deep LFG Grammars1
Researchers at PARC were pioneers in developing finite-state methods for applications in computational linguistics, and one of the original motivations was to provide a coherent architecture for the integration of lower-level lexical processing with higher-level syntactic analysis (Kaplan and Kay, 1981; Karttunen et al., 1992; Kaplan and Kay, 1994). Finite-state methods for tokenizing and morph...
متن کاملA Broad-coverage, Representationally Minimal Lfg Parser: Chunks and F-structures Are Sufficient
Amajor reason why LFG employs c-structure is because it is context-free. According to Tree-Adjoining Grammar (TAG), the only context-sensitive operation that is needed to express natural language is Adjoining, from which LFG functional uncertainty has been shown to follow. Functional uncertainty, which is expressed on the level of f-structure, would then be the only extension needed to an other...
متن کاملDeriving an Lfg from a Treebank Resource
High quality training corpora are crucial for statistical approaches to natural language processing. For probabilistic Lexical Functional Grammars (LFG-DOP, (Bod R. & Kaplan R. 1998)) significant corpora of texts associated with both c-structure and f-structure representations are required. This poses an important acquisition problem: manual construction is time-consuming and errorprone while s...
متن کاملThe Creation of a Large-Scale LFG-Based Gold Parsebank
Systems for syntactically parsing sentences have long been recognized as a priority in Natural Language Processing. Statistics-based systems require large amounts of high quality syntactically parsed data. Using the XLE toolkit developed at PARC and the LFG Parsebanker interface developed at Bergen, the Parsebank Project at Powerset has generated a rapidly increasing volume of syntactically par...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2003